Papers with LLM pretraining

9 papers
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning (2026.eacl-long)

Copied to clipboard

Challenge: Curriculum learning has improved efficiency across machine learning domains, but remains underexplored for language model pretraining.
Approach: They present a systematic investigation of curriculum learning in LLM pretraining . they use vanilla curriculum learning, pacing-based sampling, and interleaved curricula .
Outcome: The proposed framework accelerates convergence in early and mid-training phases, reducing training steps by 18-45% to reach baseline performance.
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies on detecting pretraining data in large language models have focused on sentence-level membership inference attacks (MIAs) but these methods often exhibit poor accuracy, failing to account for the semantic importance of textual content and word significance.
Approach: They propose a method that leverages established natural language processing techniques to tag keywords in input text and then uses them to obtain probabilities and calculate their average log-likelihood to determine input text membership.
Outcome: The proposed method exploits established natural language processing techniques to tag keywords in input text and calculate their average log-likelihood to determine input text membership.
Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models have demonstrated great performance across various benchmarks, but data contamination is a concern in their evaluation.
Approach: They analyze 50 papers on data contamination detection and test three of them as case studies to identify the possibility of data contamination.
Outcome: The proposed methods can detect membership inference attacks on instance-level data, and can perform similar to random guessing on LLM pretraining datasets.
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark.
Approach: They propose a benchmark for evaluating large language models on a sufficient set of scientific discovery sub-tasks.
Outcome: The proposed framework extracts critical components from papers across 12 disciplines with expert validation confirming its accuracy.
Craw4LLM: Efficient Web Crawling for LLM Pretraining (2025.findings-acl)

Copied to clipboard

Challenge: Existing work discards over 90% of the raw data collected from web crawls, highlighting the inefficiency of current web crawlers in collecting LLM pretraining data.
Approach: They propose a web crawling method that leverages the preference of LLMs as the priority score of the web crawler’s scheduler to obtain high-quality pretraining data.
Outcome: The proposed method achieves high-quality pretraining data on a web graph containing 900 million webpages from a commercial search engine's index with just 21% URLs crawled.
Proving membership in LLM pretraining data via data watermarks (2024.findings-acl)

Copied to clipboard

Challenge: Detecting whether copyright holders’ works were used in large language model (LLM) training is poised to be an important problem.
Approach: They propose to use data watermarks to enable principled detection with only black-box model access, provided the rightholder contributed multiple training documents and watermarked them before public release.
Outcome: The proposed method can be used to test hypothesis testing on a black-box model . it shows that the watermarks are strong under model and dataset scaling .
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to train LLMs consume memory equivalent to twice the model size, resulting in a hybrid design that reverts to AdamW and negates the memory gains.
Approach: They propose a new, memory-efficient O(d) adaptive scale that replaces AdamW in a hybrid structure that combines a Lion-style update direction with a memory-saving adaptive scale.
Outcome: The proposed model outperforms existing methods on LLMs up to 1.3B parameters while significantly reducing optimizer state memory.
Temporal Scaling Law for Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have found that the test loss of LLMs scales as power-laws with model size, computational budget, and dataset size.
Approach: They propose a concept of Temporal Scaling Law to study test loss of LLMs . they break down test loss into fine-grained token positions and develop a dynamic hyperbolic-law .
Outcome: The proposed model predicts the test loss of LLMs as the training steps scale up.
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction (2026.findings-acl)

Copied to clipboard

Challenge: Muon’s orthogonalized update rule suppresses the emergence of heavy-tailed weight spectra and over-emphasizes the training along noise-dominated directions.
Approach: They propose to preserve Muon's ability to capture parameter interdependencies while producing heavier-tailed updates and inducing heavier-tail weight spectra.
Outcome: The proposed algorithm suppresses the emergence of heavy-tailed weight spectra and over-emphasizes training along noise-dominated directions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations